training data
Timnit Gebru Believes There Is No 'Existential Threat' From AI
Timnit Gebru Believes There Is No'Existential Threat' From AI One of AI's fiercest critics believes the doom talk is about founders making money, not saving humanity. As the industry booms, an entirely new vernacular has emerged in the debates over how consequential artificial intelligence really is. Within that new parlance, two phrases have become focal points: stochastic parrots and existential risk . And at the center of is one prominent AI researcher who has never stood down from a fight, Timnit Gebru . Gebru came into the public eye several years ago after sparring with Google over a research paper she coauthored that called out biases in the company's AI, saying that LLMs basically parroted their training data and risked perpetuating biased viewpoints. The contested paper resulted in Gebru's departure from the company, and led her to found an institute that investigates harms perpetuated by technology and supports the creation of unbiased tech tools. She has also authored a new book,, expected to ship early next year. More recently, Gebru has spoken out against a faction of the industry that believes AI is so powerful it could destroy humanity. In turn, some members of that community, including an Anthropic cofounder, have alleged that Gebru's earlier research about stochastic parrots is no longer relevant, and that AI does have the ability to "think" or reason. The AI debate is no longer just about the technology itself, but about ideological groups and strategic narratives, and Gebru believes these narratives are a "harmful distraction" from the real issues with AI. I recently spoke with Gebru about what she believes these arguments are distracting from, and also got her response to critiques that her earlier research underestimates the AI of today.
When AI art has no author: Study finds generated images often can't be traced to training data
When AI art has no author: Study finds generated images often can't be traced to training data When an artificial intelligence image generator produces a portrait, whose work went into it? The question sits at the center of lawsuits, licensing deals, and proposed regulations worldwide. Policymakers want a way to assign responsibility. New work from a team of researchers at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) suggests that for models trained on large datasets, the question may often have no answer. It's not that the tools for finding it are inadequate.
'Scary': how misinformation and AI hallucinations are infiltrating Australia's parliament
In some cases, committee reports have cited submissions in which the majority of sources appear to be AI-generated - because they do not exist. In some cases, committee reports have cited submissions in which the majority of sources appear to be AI-generated - because they do not exist. 'Scary': how misinformation and AI hallucinations are infiltrating Australia's parliament Mon 31 Aug 2026 11.00 EDTLast modified on Mon 31 Aug 2026 11.02 EDT Australia's government inquiry process is supposed to help parliament make better decisions, hearing from experts and constituents alike. But Guardian Australia can reveal that the system is being flooded with AI-generated material, which is inventing studies and attributing nonexistent research to real academics and authors. In some cases, committee reports have cited submissions in which the majority of sources appear to be AI-generated "hallucinations", when large language models (LLMs) invent content that looks real but doesn't actually exist.
Fake US thinktank set up and funded by Israel sought to game AI for propaganda
A pro-Israel messaging website badged with the name of a thinktank that does not exist has published more than half a million words in nine days, built on a commercial platform that promises to optimize content so that AI chatbots will cite it. The site gives Israel's position on subjects including the torture of Palestinian prisoners, Israeli war crimes and whether Israel has deliberately starved Palestinians in Gaza, all presented as neutral research. The site publishes in the name of the Hanover Institute for Public Policy, which apparently does not exist as a legal entity in any jurisdiction, has no physical address, and carries no named staff and no bylines on any of its reports. Its own terms of use are governed by the laws of "the state in which the Institute is established", which they decline to name. The existence of the site was first reported on 14 August.
Japan to require AI firms to disclose training data
Japan is considering setting out a nonbinding code for generative artificial intelligence businesses to encourage them to disclose their AI training data and methods for collecting such data to the public. A government panel Tuesday broadly approved a plan to adopt what is known as a "principle code" for generative artificial intelligence businesses to protect intellectual property rights by urging firms to disclose their AI training data and methods for collecting such data to the public. A draft code was presented at an online meeting of an expert panel on intellectual property rights in the AI era. It is based on a law on AI-related technology enacted in May 2025 and seeks to balance rights protection with technological innovation. The government will use a "comply or explain" approach, under which it will set out a nonbinding code for generative AI businesses, including system developers and service providers, allowing them to choose either to comply with the code or publicly explain why they will not comply. Firms that decide to comply with the code will announce their compliance on their websites and notify the government.
Someone Is Mysteriously Snapping Up Used Books Around the World
Are AI companies to blame? Last week, the internet raged as social-media posts and news articles accused AI companies of "destroying the world's books" --including "millions of rare" ones--by chopping off their spines to scan them more easily, and discarding them afterward. One article said the news was "sparking concerns that the last remaining copies of out-of-print texts are being destroyed." The investor and AI skeptic Michael Burry called the practice of sacrificing rare books "evil incarnate." Even Elon Musk weighed in, posting that he had asked the engineers training xAI's models to "preserve any rare books in a library and scan them the hard way."
The Most Famous AI Writing Tic Is Also the Most Mysterious
If had debuted this year, William Shakespeare might have been accused of writing it with AI. A certain suspicious rhetorical device appears again and again in the play. It's in Act I, Scene ii: "The fault, dear Brutus, is not in our stars, but in ourselves." In Act III, Scene ii: "Not that I loved Caesar less, but that I loved Rome more." And later in that same scene: "I come to bury Caesar, not to praise him."
Interview with Thi Kieu Khanh Ho: Time-series anomaly detection
The latest interview in our series with the AAAI/SIGAI Doctoral Consortium participants features Thi Kieu Khanh Ho who is studying time-series anomaly detection. We found out more about her research, and what inspired her to study AI, and what she plans to work on next. Tell us a bit about your PhD -- where are you studying, and what is the topic of your research? I am doing my PhD at McGill University and Mila - Québec AI Institute, in the Department of Electrical and Computer Engineering, supervised by Professor Narges Armanfard. My research focuses on time-series anomaly detection, the problem of teaching AI systems to recognize when something unusual or abnormal is happening in complex, real-world data streams, without relying on large amounts of labeled examples.
IF-Guide: Influence Function-Guided Detoxification of LLMs
We study how training data contributes to the emergence of toxic behaviors in large language models. Most prior work on reducing model toxicity adopts *reactive* approaches, such as fine-tuning pre-trained (and potentially toxic) models to align them with human values. In contrast, we propose a *proactive* approach--IF-Guide--that leverages influence functions to identify and suppress harmful tokens in the training data. To this end, we first show that standard influence functions are ineffective at discovering harmful training records. We then present a novel adaptation that measures token-level attributions from training data to model toxicity, along with techniques for selecting toxic training documents and a learning objective that can be integrated into both pre-training and fine-tuning. Moreover, IF-Guide does not rely on human-preference data, which is typically required by existing alignment methods. In our evaluation, we demonstrate that IF-Guide substantially reduces both explicit and implicit toxicity--by up to 10$\times$ compared to uncensored models, and up to 3$\times$ compared to baseline alignment methods such as DPO and RAD--across both pre-training and fine-tuning scenarios. IF-Guide is computationally efficient: a billion-parameter model is *not necessary* for computing influence scores; a million-parameter model--with 7.5$\times$ fewer parameters--can effectively serve as a proxy for identifying harmful data.
Gaussian Mean Field Variational Inference can Overestimate Predictive Variance
Odgers, James, Riegler, Ben, Swaroop, Siddharth, Fortuin, Vincent
Mean Field Variational Inference (MFVI) is widely understood to underestimate posterior variance. By analysing conjugate Bayesian Linear Regression (BLR), we show that this characterization is incomplete: while MFVI underestimates the variance in parameter space, it can overestimate the predictive variance compared to the exact posterior. We show that if the MFVI posterior underestimates predictive variances in some directions, it necessarily overestimates them in others. Crucially, this overestimation occurs in directions where the training data concentrates. This leads to the surprising result that, for a test point drawn from the training distribution, MFVI's expected predictive variance exceeds that of the exact posterior. We demonstrate a pathological case of this effect, where the MFVI posterior fails to reduce predictive variance compared to the prior on in distribution data. We connect these results to the Cold Posterior Effect, arguing that varying the temperature can correct this overestimation, yielding predictions closer to those of the exact posterior. We validate our theory on synthetic and real-world regression tasks.